Papers with automatic generation
“Knowledge is Power”: Constructing Knowledge Graph of Abdominal Organs and Using Them for Automatic Radiology Report Generation (2023.acl-industry)
Copied to clipboard
Kaveri Kale, Pushpak Bhattacharyya, Aditya Shetty, Milind Gune, Kush Shrivastava, Rustom Lawyer, Spriha Biswas
| Challenge: | conventional radiology workflows involve dictating diagnosis to transcriptionists, which is prone to delay and error. |
| Approach: | They propose to generate a set of knowledge graphs from a large collection of free-text radiology reports and use them to generate automatic radiology report generation. |
| Outcome: | The proposed model improves the reported BLEU-3, ROUGE-L, METEOR, and CIDEr scores by 2%, 4%, 2% and 2% respectively. |
Towards Generation and Recognition of Humorous Texts in Portuguese (2023.eacl-srw)
Copied to clipboard
| Challenge: | This PhD thesis focuses on the automatic generation and recognition of verbal punning humor in Portuguese. |
| Approach: | They propose to combine natural language generation and cognitive processing to generate and recognize verbal humor in Portuguese. |
| Outcome: | The proposed methods aim to generate and recognize humor in Portuguese, an underdeveloped language compared to English. |
DIVKNOWQA: Assessing the Reasoning Ability of LLMs via Open-Domain Question Answering over Knowledge Base and Text (2024.findings-naacl)
Copied to clipboard
| Challenge: | Retrievalaugmented LLMs have been used to ground LLM in external knowledge . a gap exists in the current landscape regarding the effectiveness of grounding LLM on heterogeneous knowledge sources. |
| Approach: | They propose a model that uses symbolic language to generate symbolic queries . they use a dataset that is generated using predefined reasoning chains and human annotation . |
| Outcome: | The proposed model outperforms previous approaches by a significant margin in QA tasks over text. |
Automatic Generation of High Quality CCGbanks for Parser Domain Adaptation (P19-1)
Copied to clipboard
| Challenge: | Existing methods for Combinatory Categorial Grammar (CCG) parsing are limited to a specific parser architecture, making it non-trivial to apply to current parsers. |
| Approach: | They propose a domain adaptation method for Combinatory Categorial Grammar (CCG) they propose to generate CCG corpora using cheaper dependency trees. |
| Outcome: | The proposed method improves on speech conversation and math problems. |
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Recent studies have demonstrated that LLMs are highly sensitive to small, meaning-preserving variations in task formulation. |
| Approach: | They propose a framework that enables the automatic generation of various prompts. |
| Outcome: | The proposed framework provides meaningful variations to support strong evaluation practices. |
Assistive Recipe Editing through Critiquing (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for generating recipes that satisfy dietary restrictions are inconsistent or incoherent and paired datasets are not available at scale. |
| Approach: | They propose to build a hierarchical denoising auto-encoder that edits recipes given ingredient-level critiques by interacting with the predicted ingredients. |
| Outcome: | The proposed model can more effectively edit recipes compared to strong language models and iteratively rewrites recipes to satisfy user feedback. |
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts (2026.tacl-1)
Copied to clipboard
| Challenge: | Evaluating natural language generation systems is challenging due to the diversity of valid outputs. |
| Approach: | They propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions. |
| Outcome: | The proposed method requires only a single evaluation sample and eliminates manual prompt engineering. |
LM-BFF-MS: Improving Few-Shot Fine-tuning of Language Models based on Multiple Soft Demonstration Memory (2022.acl-short)
Copied to clipboard
| Challenge: | LM-BFF (CITATION) achieves significant few-shot performance by using auto-generated prompts and adding demonstrations similar to an input example. |
| Approach: | They propose to use auto-generated prompts and add demonstrations to LM-BFF to improve few-shot fine-tuning of language models with multiple soft demonstrations. |
| Outcome: | The proposed method improves few-shot fine-tuning on eight NLP tasks. |
EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models (2025.emnlp-demos)
Copied to clipboard
Ziwen Xu, Shuxun Wang, Kewei Xu, Haoming Xu, Mengru Wang, Xinle Deng, Yunzhi Yao, Guozhou Zheng, Huajun Chen, Ningyu Zhang
| Challenge: | Large Language Models (LLMs) have demonstrated extraordinary capabilities, however, they may still generate unreliable or unsafe outputs. |
| Approach: | They propose a framework that allows plug-and-play adjustability for controlling Large Language Model (LLM) behaviors. |
| Outcome: | The framework is designed to enable plug-and-play adjustability for controlling Large Language Model (LLM) behaviors. |
Generating Vehicular Icon Descriptions and Indications Using Large Vision-Language Models (2024.emnlp-industry)
Copied to clipboard
James Fletcher, Nicholas Dehnen, Seyed Nima Tayarani Bathaie, Aijun An, Heidar Davoudi, Ron DiCarlantonio, Gary Farmaner
| Challenge: | Existing image description systems are trained mainly on natural images, whereas icon images are drawings. |
| Approach: | They propose to use a dataset to generate both visual and functional icon descriptions based on the icon image and its context information in the car manual. |
| Outcome: | The proposed model performs well on the dashboard icon description task while the third model perform poorly. |
Towards Continuous Dialogue Corpus Creation: writing to corpus and generating from it (L18-1)
Copied to clipboard
| Challenge: | Existing methods to create dialogue corpora annotated with interoperable semantic information are based on ISO standard data models and tools. |
| Approach: | They propose to use a corpus as a shared repository for analysis and modelling of interactive dialogue behaviour and for implementation, integration and evaluation of dialogue system components. |
| Outcome: | The proposed method is applied to the design of two multimodal interactive applications - the Virtual Negotiation Coach and the Virtual Debate Coach. |
Self-Contained Utterance Description Corpus for Japanese Dialog (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing task frameworks for dialog-act classification and slot filling can only interpret utterances using pre-defined types and slots. |
| Approach: | They propose a task to describe the intent of an utterance in a dialog with multiple simple natural sentences without the context. |
| Outcome: | The proposed task can describe the intent of an utterance in a dialog with multiple simple natural sentences without the context. |
SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Automated survey generation is a key task in scientific document processing due to lack of standardized evaluation datasets. |
| Approach: | They propose a survey-based framework that integrates quality indicators into literature retrieval to assess higher-quality sources. |
| Outcome: | The proposed framework enhances the standard Retrieval-Augmented Generation pipeline and enables human-guided writing. |
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have gained increasing attention for their capacity to generate harmful content. |
| Approach: | They propose a scalable evolution framework to evolve red teaming prompts across breadth and depth dimensions, facilitating automatic generation of numerous high-quality and diverse red team prompts. |
| Outcome: | The proposed framework surpasses existing red teaming methods on attack success rate and diversity. |
The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents (2023.eacl-main)
Copied to clipboard
| Challenge: | StatCan Dialogue Dataset consists of 19,379 conversation turns between agents and online users . researchers propose two tasks to help knowledge workers find relevant tables for live chat users based on real-world intents . |
| Approach: | They propose two tasks based on 19,379 conversation turns between agents and online users . they investigate the difficulty of each task by establishing strong baselines . |
| Outcome: | The proposed task is based on a dataset of 19,379 conversation turns . the researchers show that the models struggle to generalize to future conversations . |
Cross-modal Contrastive Attention Model for Medical Report Generation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for medical report generation are unable to capture useful information from historical cases. |
| Approach: | They propose a model that captures both visual and semantic information from similar cases. |
| Outcome: | The proposed model outperforms the state-of-the-art models on almost all metrics on IU X-Ray and MIMIC-CXR benchmarks. |
On the Automatic Generation and Simplification of Children’s Stories (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have made it possible to generate children's educational texts with appropriate lexical and readability levels. |
| Approach: | They first examine the ability of several popular LLMs to generate stories with properly adjusted lexical and readability levels. |
| Outcome: | The proposed models can generalize to the domain of children's stories and create an efficient pipeline for their automatic generation. |
One Comment from One Perspective: An Effective Strategy for Enhancing Automatic Music Comment (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for automatic comment generation generate common and meaningless comments for music. |
| Approach: | They propose a multi-perspective strategy to enhance automatic music comment generation by combining different perspectives on a music comment dataset. |
| Outcome: | The proposed model outperforms state-of-the-art models on two music comment datasets and outperformed existing models by a substantial margin. |
ChainLM: Empowering Large Language Models with Improved Chain-of-Thought Prompting (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing CoT synthesis approaches focus on simpler reasoning tasks and result in inconsistent CoT prompts. |
| Approach: | They propose a framework for automatic generation of superior CoT prompts based on three major evolution strategies . they propose 'step-level debating' method where multiple debaters discuss each reasoning step to arrive at the correct answer. |
| Outcome: | The proposed framework can generate superior CoT prompts from a CoT dataset. |
Constructing Highly Inductive Contexts for Dialogue Safety through Controllable Reverse Generation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to detect toxic generation of pretrained language models rely on templates, data extraction, crowdsourcing workers or automatic generation. |
| Approach: | They propose a method to construct adversarial contexts conditioned on a given response . they augment existing dataset BAD+ and construct a new dataset B AD+ . |
| Outcome: | The proposed method can detect toxic or biased content in large pretrained language models. |
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Literature review tables are essential for summarizing and comparing collections of scientific papers. |
| Approach: | They propose to generate a database of literature review tables from a pool of papers and to model retrieval noise via semantically related but out-of-scope distractor papers verified by human annotators. |
| Outcome: | The proposed method improves over strong baselines while the absolute scores remain modest, underscoring the task’s difficulty. |
Generation of a Spanish Artificial Collocation Error Corpus (L18-1)
Copied to clipboard
| Challenge: | collocations are combinations of two elements where one (the base) is freely chosen, despite the limitations of the other (collocate) current tools for collocation error detection and correction focus on collocation validation and identification of miscollocations . |
| Approach: | They propose an algorithm for automatic generation of an artificial collocation error corpus of american English learners of Spanish that includes 17 different types of collocation errors. |
| Outcome: | The proposed algorithm can detect and classify collocation errors in learners' writings . collocation error detection and correction has not received the attention it deserves . |
PAP2PAT: Benchmarking Outline-Guided Long-Text Patent Generation with Patent-Paper Pairs (2025.findings-acl)
Copied to clipboard
| Challenge: | In patents, the description constitutes more than 90% of the document on average, yet its automatic generation remains understudied. |
| Approach: | They propose a method to generate patent documents using a research paper as an invention specification. |
| Outcome: | The proposed model can generate 1.8k patent-paper pairs describing the same inventions, but it's difficult to provide the level of detail required. |
Multi-Figurative Language Generation (2022.coling-1)
Copied to clipboard
| Challenge: | Figurative language generation is the task of reformulating a given text in the desired figure of speech while still being faithful to the original context. |
| Approach: | They propose a scheme for multi-figurative language pre-training on top of BART and a mechanism for injecting the target figurative information into the encoder to generate text with the target figure from another figurativ form without parallel figura-figura pairs. |
| Outcome: | The proposed model outperforms all baselines and qualitatively examines the relationship between the different figures of speech. |
Generation of Patient After-Visit Summaries to Support Physicians (2022.coling-1)
Copied to clipboard
Pengshan Cai, Fei Liu, Adarsha Bajracharya, Joe Sills, Alok Kapoor, Weisong Liu, Dan Berlowitz, David Levy, Richeek Pradhan, Hong Yu
| Challenge: | After-visit summary is a summary note given to patients after their clinical visit. |
| Approach: | They propose to automate the generation of after-visit summaries and introduce a feedback mechanism that alerts physicians when an automatic summary fails to capture important details of the clinical notes. |
| Outcome: | The proposed system improves on a large clinical dataset that contains electronic health record (EHR) notes and their associated summaries. |
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to eliminate implicit biases in LLMs do not eradicate underlying behavioral bias. |
| Approach: | They propose a framework that uses logic grid puzzles to probe the influence of social stereotypes on logical reasoning and decision making in LLMs. |
| Outcome: | The proposed framework systematically probes the influence of social stereotypes on logical reasoning and decision making in LLMs. |
A Symbolic Adversarial Learning Framework for Evolving Fake News Generation and Detection (2025.emnlp-main)
Copied to clipboard
| Challenge: | Rapid LLM advancements heighten fake news risks by enabling the automatic generation of increasingly sophisticated misinformation. |
| Approach: | They propose a framework that implements an adversarial training paradigm by an agent symbolic learning optimization process rather than numerical updates. |
| Outcome: | The proposed framework generates sophisticated fake news that degrades state-of-the-art detection performance by 53.4% in Chinese and 34.2% in English on average. |
Interpretable Multimodal Misinformation Detection with Logic Reasoning (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for misinformation detection lack interpretability due to the black-box nature of the neural network. |
| Approach: | They propose a logic-based neural model which integrates interpretable logic clauses to express the reasoning process of the target task. |
| Outcome: | The proposed model can be generalizable across multiple misinformation sources and is based on three public datasets. |
Improve LLM-as-a-Judge Ability as a General Ability (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies focus on generative judges, but only on their judge ability. |
| Approach: | They propose a method that leverages the generative and reasoning capabilities of large language models to evaluate LLM responses across diverse scenarios, providing accurate preference signals. |
| Outcome: | The proposed model performs on RewardBench with only 2% to 40% of the data required by other training frameworks. |
Evaluating the Knowledge Dependency of Questions (2022.emnlp-main)
Copied to clipboard
Hyeongdon Moon, Yoonseok Yang, Hangyeol Yu, Seunghyun Lee, Myeongho Jeong, Juneyoung Park, Jamin Shin, Minsam Kim, Seungtaek Choi
| Challenge: | Existing evaluation metrics for MCQ generation focus on the n-gram based similarity of the generated MCq to the gold sample and disregard their educational value. |
| Approach: | They propose to use a human survey to measure the MCQ’s answerability given knowledge of the target fact. |
| Outcome: | The proposed methods measure the MCQ’s answerability given knowledge of the target fact. |
Annotating a Fable in Italian Sign Language (LIS) (2020.lrec-1)
Copied to clipboard
| Challenge: | fables are short or medium-length stories with a moral and they generally have specific characteristics in SLs that are usually not to be found in spoken languages like Italian. |
| Approach: | They present work for automatic generation of a written text in Italian starting from glosses of fable in Italian Sign Language (LIS). |
| Outcome: | The proposed method was used to generate a written text in Italian starting from glosses of a fable in Italian Sign Language (LIS). |
Robustifying Sentiment Classification by Maximally Exploiting Few Counterfactuals (2022.emnlp-main)
Copied to clipboard
| Challenge: | a recent study found that finetuned language models rely on spurious patterns in training data . this limitation limits their performance on out-of-distribution (OOD) test data. |
| Approach: | They propose a method that only requires annotation of a small fraction of training data . they add 1% manual counterfactuals to training data and generate extra counterfacts in vector space . |
| Outcome: | The proposed approach improves sentiment classification using IMDb data and other sets for OOD tests. |
A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation Discovery (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have proposed to take advantage of the scientific paper's citation network to approach literature summarization. |
| Approach: | They propose to annotate related work sections, cite papers and sentences using machine readable data and an additional layer of papers citing the references. |
| Outcome: | The proposed corpus expands the existing data-set of related work sections and cites the papers cited in the related work section. |
GTA: Generating Long-horizon Tasks for Web Agents at Scale (2026.acl-long)
Copied to clipboard
Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, Chien-Sheng Wu
| Challenge: | Existing benchmarks provide only coarse start–goal annotations without intermediate trajectories . Existing frameworks provide no supervision over the agent's latent decision process . |
| Approach: | They propose a framework that integrates crawling, retrieval-based seeding, in-context generation and automated quality control to produce realistic tasks paired with executable trajectories. |
| Outcome: | The proposed framework decouples crawling from generation for greater efficiency and ensures dense supervision through deterministic replays and systematic validation. |
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies on proactive dialogue models focus on domain-specific or task-oriented scenarios, which leads to fragmented evaluations and limits the comprehensive exploration of models’ proactive dialogue abilities. |
| Approach: | They propose a framework for evaluating proactive dialogue capabilities of large language models that decomposes proactive dialogue into target planning and dialogue guidance, establishing evaluation metrics across various domains. |
| Outcome: | The proposed framework decomposes proactive dialogue into target planning and dialogue guidance, establishing evaluation metrics across various domains, and enables automatic generation of diverse and challenging evaluation data. |